Papers with dense representations
Pretrained Transformers for Text Ranking: BERT and Beyond (2021.naacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial provides an overview of text ranking using neural network architectures known as transformers. |
| Approach: | This tutorial provides an overview of text ranking with neural network architectures known as transformers. |
| Outcome: | This tutorial provides an overview of text ranking with neural network architectures known as transformers. |
Extracting Text Representations for Terms and Phrases in Technical Domains (2023.acl-industry)
Copied to clipboard
| Challenge: | Large pre-trained language models are extensively used in modern NLP systems. |
| Approach: | They propose an unsupervised approach to encoding using character-based models and pre-trained sentence encoders to reconstruct large pre-trained embedding matrices. |
| Outcome: | The proposed approach matches the quality of sentence encoders in technical domains and is 5 times smaller and up to 10 times faster on high-end GPUs. |
GNN-encoder: Learning a Dual-encoder Architecture via Graph Neural Networks for Dense Passage Retrieval (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to perform large-scale query-passage retrieval are term-based, but they lose interaction between query-pastage pairs. |
| Approach: | They propose to fuse query (passage) information into query representations via graph neural networks that are constructed by queries and their top retrieved passages. |
| Outcome: | The proposed model outperforms existing models on MSMARCO, Natural Questions and TriviaQA datasets and achieves the new state-of-the-art on these datasets. |
Contextualized Query Embeddings for Conversational Search (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to conversational search use multiple inference pipelines that require long inference times . despite their effectiveness, such a pipeline often includes multiple neural models that require longer inference time. |
| Approach: | They propose to integrate conversational query reformulation directly into a dense retrieval model . they use a dataset with pseudo-relevance labels to overcome the lack of training data . |
| Outcome: | The proposed model rewrites conversational queries as dense representations in conversational search and open-domain question answering datasets. |
The Curse of Dense Low-Dimensional Information Retrieval for Large Index Sizes (2021.acl-short)
Copied to clipboard
| Challenge: | Existing studies have shown that dense representations outperform sparse representations with large index sizes. |
| Approach: | They propose to use dense low-dimensional representations to retrieve relevant documents . they show performance decreases quicker for increasing index sizes than for sparse representations . |
| Outcome: | The proposed representations outperform sparse representations with large index sizes. |
Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text Retrieval (2021.emnlp-main)
Copied to clipboard
| Challenge: | Recent approaches to information retrieval (IR) and natural language processing (NLP) use contextual language models, which can improve both synonymy and polysemy problems associated with words. |
| Approach: | They propose an ultra-high dimensional representation scheme equipped with directly controllable sparsity and a bucketing method where embeddings from multiple layers of BERT are selected/merged to represent diverse linguistic aspects. |
| Outcome: | The proposed representation scheme outperforms sparse models with MS MARCO and TREC CAR, and shows that it is highly efficient for storage and search. |
Explaining Relationships Between Scientific Documents (2021.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to explain relationships between scientific documents using natural language text can be useful for research efficiency. |
| Approach: | They propose a task of explaining relationships between scientific documents using natural language text. |
| Outcome: | The proposed models can be automated and humanely evaluated. |
Learning Target-Specific Representations of Financial News Documents For Cumulative Abnormal Return Prediction (C18-1)
Copied to clipboard
| Challenge: | Recent work considers learning dense representations for news titles and abstracts . text representations can address the sparsity of discrete indicators in statistical models . |
| Approach: | They propose to use news abstracts to combine the most informative sentences in news content to learn dense representations for text elements. |
| Outcome: | The proposed model can be used to estimate abnormal returns of companies when compared to titles and abstracts. |
Transformation of Dense and Sparse Text Representations (2020.coling-main)
Copied to clipboard
| Challenge: | Existing approaches to NLP to leverage sparsity have been limited due to the gap with dense representations. |
| Approach: | They propose a Semantic Transformation method to bridge dense and sparse spaces and propose supervised NLP tasks to use both spaces. |
| Outcome: | Experiments with classification tasks and natural language inference tasks show that the proposed method is effective. |
Multi-Step Reasoning Over Unstructured Text with Beam Dense Retrieval (2021.naacl-main)
Copied to clipboard
| Challenge: | Current methods for complex question answering use structured knowledge and unstructured text. |
| Approach: | They propose a multi-step retrieval approach that iteratively forms an evidence chain through beam search in dense representations. |
| Outcome: | The proposed method is competitive to state-of-the-art systems without using semi-structured information. |
Unifying Multimodal Retrieval via Document Screenshot Embedding (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing document retrieval pipelines require document parsing and content extraction to prepare input for indexing. |
| Approach: | They propose a retrieval paradigm that regards document screenshots as a unified input format . they leverage a large vision-language model to directly encode document screenshot into dense representations . |
| Outcome: | The proposed method outperforms existing retrieval pipelines in a text-intensive context. |
Improving Document Representations by Generating Pseudo Query Embeddings for Dense Retrieval (2021.acl-long)
Copied to clipboard
| Challenge: | Existing retrieval models based on dense representations show better performance than sparse representations. |
| Approach: | They propose a method to mimic the queries to each of the documents by an iterative clustering process and represent the documents using multiple pseudo queries. |
| Outcome: | The proposed model achieves state-of-the-art results on a large dataset while remaining high efficiency. |
Natural Logic-guided Autoregressive Multi-hop Document Retrieval for Fact Verification (2022.emnlp-main)
Copied to clipboard
| Challenge: | Recent evidence retrieval approaches rely on heuristics and assume hyperlinks between documents. |
| Approach: | They propose a retrieval method that combines a retriever and a proof system that reranks documents and reorders them . |
| Outcome: | The proposed method exceeds or is on par with the current state-of-the-art on FEVER, HoVer and FEVEROUS-S while using 5 to 10 times less memory than competing systems. |
Learning Dense Representations of Phrases at Scale (2021.acl-long)
Copied to clipboard
| Challenge: | Existing phrase retrieval models rely on sparse representations and still underperform retriever-reader approaches. |
| Approach: | They propose a method to learn phrase representations from reading comprehension tasks using negative sampling methods. |
| Outcome: | The proposed model improves over previous models by 15%-25% absolute accuracy and matches the performance of state-of-the-art retrieval models. |
Dense Passage Retrieval for Open-Domain Question Answering (2020.emnlp-main)
Copied to clipboard
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, Wen-tau Yih
| Challenge: | Open-domain question answering relies on efficient passage retrieval to select candidate contexts. |
| Approach: | They propose a dual-encoder framework that can be implemented to retrieve passages from a small number of questions and passages. |
| Outcome: | The proposed system outperforms a strong Lucene-BM25 system in top-20 passage retrieval accuracy on multiple open-domain QA benchmarks. |
MIST: Mutual Information Maximization for Short Text Clustering (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods for clustering short texts are inadequate due to the limited amount of information provided by each text sample. |
| Approach: | They propose a Mutual Information Maximization Framework for Short Text Clustering which maximizes mutual information between representations on sequence and token levels. |
| Outcome: | The proposed framework outperforms the state-of-the-art method in terms of Accuracy or Normalized Mutual Information in most cases. |
STAIR: Learning Sparse Text and Image Representation in Grounded Tokens (2023.emnlp-main)
Copied to clipboard
Chen Chen, Bowen Zhang, Liangliang Cao, Jiguang Shen, Tom Gunter, Albin Jose, Alexander Toshev, Yantao Zheng, Jonathon Shlens, Ruoming Pang, Yinfei Yang
| Challenge: | State-of-the-art contrastive learning models like CLIP and ALIGN are less interpretable and suffer from inferior accuracy than dense representations. |
| Approach: | They extend CLIP and ALIGN models to build a sparse semantic representation that is interpretable and easy to integrate with existing retrieval systems. |
| Outcome: | The proposed model outperforms CLIP and ALIGN models on image and text retrieval tasks with a 4.9% and +4.3% improvement on COCO-5k textimage and imagetext retrieval respectively. |
ConvX: A Lightweight Converter to Bridge Indexed Dense Representations and Large Language Models for Retrieval-Augmented Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing RAG pipelines suffer from critical efficiency limitations due to their complexity and complexity. |
| Approach: | They propose a compression-based RAG framework that directly leverages indexed dense representations produced by a retriever, substituting to long text contexts. |
| Outcome: | Empirical results show that the proposed model achieves competitive performances compared to the state-of-the-art model that uses a large ad-hoc context compressor while offering substantially improved inference efficiency. |
Decoupled Reasoning with Implicit Fact Tokens (DRIFT): A Dual-Model Framework for Efficient Long-Context Inference (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing solutions to integrate extensive, dynamic knowledge into Large Language Models (LLMs) are constrained by finite context windows, retriever noise, or the risk of catastrophic forgetting. |
| Approach: | They propose a dual-model architecture that explicitly decouples knowledge extraction from the reasoning process by compressing document chunks into implicit fact tokens conditioned on the query. |
| Outcome: | The proposed architecture significantly outperforms strong baselines among comparably sized models on long-context tasks while maintaining inference accuracy. |